Papers by David R. Mortensen

27 papers
A Hmong Corpus with Elaborate Expression Annotations (2022.lrec-1)

Copied to clipboard

Challenge: SCH is the first substantial corpus to be annotated for elaborate expressions . a plurality of speakers are located in China, but many Hmong speakers left Laos as refugees .
Approach: They describe the first publicly available corpus of Hmong, a minority language of China, Vietnam, Laos, Thailand, and various countries in Europe and the Americas.
Outcome: The first publicly available corpus of Hmong is scraped from a long-running Usenet newsgroup . it is the first substantial corpus to be annotated for elaborate expressions .
Adapting Word Embeddings to New Languages with Morphological and Phonological Subword Representations (D18-1)

Copied to clipboard

Challenge: Existing approaches to generalization to resource-rich languages are difficult . a recent study shows that word representations can be useful in low resource languages .
Approach: They propose two approaches for improving generalization to low-resource languages by adapting continuous word representations using linguistically motivated subword units.
Outcome: The proposed method improves generalization to low resource languages . it requires neither parallel corpora nor bilingual dictionaries and requires no parallel training .
Improved Neural Protoform Reconstruction via Reflex Prediction (2024.lrec-main)

Copied to clipboard

Challenge: comparative method allows linguists to infer protoforms from their reflexes based on sound change . authors argue that this approach ignores one of the most important aspects of the comparative approach .
Approach: They propose a comparative method that allows linguists to infer protoforms from their reflexes . they propose to use a system where candidate protoform from a reconstruction model are reranked by a reflex prediction model.
Outcome: The comparative method surpasses state-of-the-art methods on Chinese and Romance datasets.
Cross-Cultural Similarity Features for Cross-Lingual Transfer Learning of Pragmatically Motivated Tasks (2021.eacl-main)

Copied to clipboard

Challenge: a large amount of work on cross-lingual transfer learning focused on typological and genealogical similarities between languages.
Approach: They propose three features that capture cross-cultural similarities that manifest in linguistic patterns and quantify distinct aspects of language pragmatics.
Outcome: The proposed features capture cross-cultural similarities manifest in linguistic patterns and quantify aspects of language pragmatics.
ZIPA: A family of efficient models for multilingual phone recognition (2025.acl-long)

Copied to clipboard

Challenge: IPA transcriptions capture major articulatory contrasts in speech sounds, including the voicing status, place of articulation, manner of voicing, and tongue positions.
Approach: They present ZIPA, a family of efficient speech models that advances the state-of-the-art performance of crosslinguistic phone recognition.
Outcome: The proposed model outperforms existing phone recognition systems on 17,000+ hours of normalized phone transcriptions and a novel evaluation set capturing unseen languages and sociophonetic variation.
WikiHan: A New Comparative Dataset for Chinese Languages (2022.coling-1)

Copied to clipboard

Challenge: Currently, there are 1.3 billion speakers of Sinitic varieties, making the family one of the largest in terms of speaker count.
Approach: They have collected a single constituent and structured form of Chinese varieties for comparative linguistics and Chinese NLP.
Outcome: The proposed dataset contains 67,943 entries across 8 varieties and Middle Chinese . it achieves 54.11% accuracy and 17.69% error rate on a protoform reconstruction task .
POWSM: A Phonetic Open Whisper-Style Speech Foundation Model (2026.acl-long)

Copied to clipboard

Challenge: Phone-level modeling of speech is a common approach to speech recognition, but it relies on task-specific architectures and datasets.
Approach: They propose a phonetic framework capable of performing multiple phone-related tasks . they propose 'Phonetic Open Whisper-style Speech Model' that can perform these tasks together .
Outcome: The proposed model outperforms or matches specialized PR models of similar size while supporting G2P, P2G, and ASR.
Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration (2026.findings-acl)

Copied to clipboard

Challenge: We show that script information is linearly encoded in the activation space of multilingual speech models . modifying activations at inference time induces script change even in unconventional pairings .
Approach: They propose to add script vectors to activations at test time to induce script change . they also show that script information is linearly encoded in the activation space of multilingual speech models .
Outcome: The proposed approach can induce script change even in unconventional language-script pairings.
Verbing Weirds Language (Models): Evaluation of English Zero-Derivation in Five LLMs (2024.lrec-main)

Copied to clipboard

Challenge: Lexical-syntactic flexibility is a hallmark of English morphology . conversion involves placing a word with one part of speech in a non-prototypical context .
Approach: They propose to test lexical-syntactic flexibility in the form of conversion . conversion is a process where a word with one part of speech is placed in a non-prototypical context .
Outcome: The proposed task tests the ability of five language models to generalize over words with a non-prototypical part of speech.
Happiness is Sharing a Vocabulary: A Study of Transliteration Methods (2026.eacl-long)

Copied to clipboard

Challenge: a key problem in multilingual NLP is script barrier, which makes it difficult to share knowledge between languages . a new study shows that transliteration can be useful for languages using non-Latin scripts .
Approach: They propose to use romanization, phonemic transcription, and substitution ciphers to evaluate models . romanization outperforms other input types in 7 out of 8 evaluation settings .
Outcome: The proposed approach outperforms other input types on three tasks and is the most effective . romanization outperformed other input type in 7 out of 8 evaluation settings .
Programming by Example meets Historical Linguistics: A Large Language Model Based Approach to Sound Law Induction (2025.acl-long)

Copied to clipboard

Challenge: Historical linguists have written programs that convert reconstructed words into their attested descendants via ordered string rewrite functions.
Approach: They propose to use a model to generate a "similar distribution" for sound law induction . they propose four kinds of methods with varying amounts of inductive bias to investigate best performance .
Outcome: The proposed model shows that it can be fine tuned with training data and evaluation data.
Automatic Extraction of Rules Governing Morphological Agreement (2020.emnlp-main)

Copied to clipboard

Challenge: Creating a descriptive grammar is an indispensable step for language documentation but it is tedious and time-consuming.
Approach: They propose a framework for extracting a first-pass grammatical specification from raw text in a concise, human- and machine-readable format.
Outcome: The proposed framework extracts a grammatical specification that is nearly equivalent to those created with large amounts of gold-standard annotated data.
Phonotactic Complexity across Dialects (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies show a moderate negative correlation between phonotactic complexity and word length in 106 languages.
Approach: They propose to use a phone-level language model to measure phonotactic complexity . they find a tradeoff between word length and phonomactic complex .
Outcome: The proposed model shows that low phonotactic complexity dialects concentrate around capital regions.
Epitran: Precision G2P for Many Languages (L18-1)

Copied to clipboard

Challenge: Epitran is a multilingual, multi-back-end system for grapheme-to-phoneme transduction . it supports 61 languages and is open source under an MIT license .
Approach: Epitran is a multilingual back-end system for grapheme-to-phoneme transduction . it takes word tokens in the orthography of a language and outputs a phonemic representation . Epitran's efficacy has been demonstrated in multiple research projects .
Outcome: Epitran is a multilingual, multi-backend system for grapheme-to-phoneme transduction . it supports 61 languages and is open source under MIT license .
Searching for the Most Human-like Emergent Language (2025.emnlp-main)

Copied to clipboard

Challenge: Existing work on emergent communication systems to generate languages with high statistical similarity to human languages has not been done.
Approach: They propose to optimize a signalling game-based emergent communication environment to generate state-of-the-art emergentic languages with a high degree of similarity to human language.
Outcome: The proposed language generates state-of-the-art on XferBench benchmark, demonstrating its similarity to human language and entropy-minimization properties.
[b] = [d] - [t] + [p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies on how self-supervised speech models encode rich phonetic information have not explored how they are structured.
Approach: They conduct a comprehensive analysis of the underlying structure of S3M representations with particular attention to phonological vectors.
Outcome: The proposed model encodes phonologically interpretable and compositional vectors, demonstrating phonology vector arithmetic.
PRiSM: Benchmarking Phone Realization in Speech Models (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluations of phone recognition systems only measure surface-level transcription accuracy.
Approach: They propose to standardize transcription-based evaluation and assess downstream utility in clinical, educational, and multilingual settings with transcription and representation probes.
Outcome: The proposed system outperforms LALMs in clinical, educational, and multilingual settings.
PWESuite: Phonetic Word Embeddings and Tasks They Facilitate (2024.lrec-main)

Copied to clipboard

Challenge: Existing word embedding methods overlook phonetic information that is crucial for many tasks.
Approach: They propose three methods that use articulatory features to build phonetically informed word embeddings.
Outcome: The proposed methods improve word retrieval and correlation with sound similarity and on rhyme and cognate detection tasks.
Evaluating the Morphosyntactic Well-formedness of Generated Texts (2021.emnlp-main)

Copied to clipboard

Challenge: Text generation systems are ubiquitous in natural language processing applications, but evaluation of these systems remains a challenge, especially in multilingual settings.
Approach: They propose a metric to evaluate the morphosyntactic well-formedness of text using its dependency parse and morphologically-rich rules of the language.
Outcome: The proposed metric can evaluate the morphosyntactic well-formedness of text using its dependency parse and morphologically-rich rules of the language.
PBEBench: A Multi-Step Programming by Examples Reasoning Benchmark inspired by Historical Linguistics (2026.findings-acl)

Copied to clipboard

Challenge: a benchmark for inductive reasoning is based on sound law induction in historical linguistics . solve rates are below 5% on hard PBEBench instances with long program cascades despite expensive scaling strategies .
Approach: They propose a benchmark for inductive reasoning inspired by sound law induction in historical linguistics.
Outcome: The proposed approach generates problems with controllable difficulty and ordering constraints . solve rates remain below 5% on hard PBEBench instances with long program cascades .
Morpheme Induction for Emergent Language (2025.emnlp-main)

Copied to clipboard

Challenge: CSAR is a greedy algorithm that weights morphemes based on mutual information between forms and meanings, then removes it from the corpus and repeats the process to induce more morphs.
Approach: They propose an algorithm that weights morphemes based on mutual information between forms and meanings, selects highest-weighted pair, removes it from corpus, and repeats process to induce further morphs.
Outcome: The proposed algorithm makes reasonable predictions in adjacent domains.
DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to Models (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in MT quality and language coverage have shown that language varieties with low baseline performance are more likely to benefit from these approaches.
Approach: They propose a training-time technique for adapting a pretrained model to dialectal data and an inference-time intervention adapting dialectal datasets to the model expertise.
Outcome: The proposed model shows significant performance gains for several dialects from four language families, and modest gains for two other language families.
Phone Inventories and Recognition for Every Language (2022.lrec-1)

Copied to clipboard

Challenge: Identifying phone inventories is crucial component in language documentation and preservation of endangered languages.
Approach: They propose a probabilistic and non-probabilistic phone inventory model that estimates the phone inventory for any language listed in Glottolog.
Outcome: The proposed model outperforms baseline models by 6.5 F1 and improves the PER (phone error rate) in phone recognition by 25%.
Transformed Protoform Reconstruction (2023.acl-short)

Copied to clipboard

Challenge: Historical linguists reconstruct proto-languages by identifying systematic sound changes that can be inferred from correspondences between attested daughter languages.
Approach: They propose to update their Latin protoform reconstruction model with the Transformer . romance data of 8,000 cognates spanning 5 languages and Chinese dataset are outperformed .
Outcome: The proposed model outperforms previous models on Romance and Chinese datasets.
Communicating in Emergent Language with an Induced Morphological Phrasebook (2026.acl-long)

Copied to clipboard

Challenge: a major challenge in studying emergent languages is interpreting how they convey meaning-neural networks may invent communication systems lacking features of human language.
Approach: They build rule-based emergent language agents using form-meaning mappings induced from ELs and test their communicative performance in the EL environment.
Outcome: The proposed model shows that EL agents rely on repetition and morpheme ordering to convey meaning.
Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons (2024.lrec-main)

Copied to clipboard

Challenge: In this paper, we examine the ability of large language models (LLMs) to identify different meanings in sentences that are superficially similar.
Approach: They propose a challenge dataset for NLP with large lexical overlap which minimises the possibility of models discerning entailment solely based on token distinctions.
Outcome: The proposed model fails to distinguish between constructions with three classes of adjectives which cannot be distinguished by surface features.
AlloVera: A Multilingual Allophone Database (2020.lrec-1)

Copied to clipboard

Challenge: Phonemes are contrastive phonological units, and allophones are their various concrete realizations.
Approach: They propose a resource that maps allophones to phonemes for 14 languages . they propose phonological representations that are much closer to a universal transcription .
Outcome: The proposed resource maps from 218 allophones to phonemes for 14 languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations